Skip to content

docs: define backend plan split and SDS state contract - #737

Merged
zzylol merged 45 commits into
mainfrom
docs/physical-plan-design
Sep 28, 2026
Merged

zzylol merged 45 commits into
mainfrom
docs/physical-plan-design

Conversation

@zzylol

@zzylol zzylol commented Sep 17, 2026 •

Copy link
Copy Markdown
Contributor

Problem and design

Before this PR, logical DAGs and Backend semantic-node bindings obscured the boundary between physical computation and deployment. Stored-state identity and recovery lacked a single design contract.

After this PR, Planner exposes supported, semantically legal physical workload candidates. Backend evaluates deployment feasibility and scoped costs, selects a candidate, and binds its unchanged Physical DAGs into PrecomputePlan and QueryPlan. Both engines use the shared physical executor.

The integration, SDS, glossary and migration documents define:

Scope and dependencies

Target design and migration requirements only; this PR does not claim these new requirements are implemented. Ad-hoc SDS discovery remains future work.

#768 is merged. Current main stack: main → #737 → #749 → #771 → #774 → #763 → #765 → #761 → #728 → #742 → #775. Diagnostics #756 and runtime inspection #766 branch from #761. #776 → #777 → #778 → #759 are deferred drafts.

Validation

Local links and heading anchors checked across the four updated documents; git diff --check passed. No runtime tests were run for this documentation-only update.

@zzylol zzylol changed the title docs: clarify Planner physical plan and SDS architecture docs: define backend plan split and SDS state contract Sep 18, 2026
## Purpose and scope

## Design decision
This design splits one selected ASAPPlanner semantic DAG into two executable

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

What is a "semantic DAG"? Is this the output of ASAPPLanner? How is this different from SummaryMaintenanceLifecyclePlan?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

“semantic DAG” means the selected post-ASAP computation DAG produced by ASAPPlanner.
SummaryMaintenanceLifecyclePlan means after a post-ASAP DAG being generated, some logic of summary maintainance will output the plan for how a summary is maintiend, incremental vs built from data at rest. So these are two steps currently in the code.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I didn't understand the relationship between Semantic DAG and SummaryMaintenanceLifecyclePlan. Also, what is ASAPQuery-backend inputting from ASAPPlanner? One of these or both?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

this relationship will be answered in ProjectASAP/ASAPPlanner#445

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The backend uses both, at different stages:

  1. Selected post-ASAP DAG: The backend’s selection adapter obtains a Rc root and passes it to physical compilation as QueryCompilationInput.selected_plan_root. This describes the
    selected computation. Code

  2. Maintenance lifecycle decisions: During physical compilation, select_lifecycle() calls Planner’s plan_summary_maintenance_lifecycles() with the selected summary node, workload demand,
    capabilities, and cost evidence. That returns a SummaryMaintenanceLifecyclePlan. The backend extracts its selected lifecycle guarantee, window framework, implementation identity, and
    costing information into backend configuration. Code

So the current flow is:

Select post-ASAP computation
→ pass selected root to backend compiler
→ compiler calls Planner for summary lifecycle decisions
→ combine computation and selected maintenance decisions
→ generate backend plans

The lifecycle plan refers to the summary computation it is planning maintenance for. It does not replace the DAG, and the backend does not simply execute the lifecycle-plan object directly.

@zzylol zzylol Sep 18, 2026 •

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Following up on my explanation above: that describes the current backend call sequence. Based on ASAPPlanner #445, I propose moving lifecycle-aware selection before physical compilation and passing its result directly to the compiler.

Current implementation

Select complete Post-ASAP DAG
    -> backend physical compiler receives selected_plan_root
         -> extracts summary producers
         -> calls Planner lifecycle API for selected producers
         -> extracts lifecycle/window decisions
         -> combines those decisions with the complete query DAG
         -> generates backend physical plans

The current compiler calls select_lifecycle(..., &selected.node, ...) on extracted producers (code). Those producer-local lifecycle results do not necessarily contain downstream query readouts, so the compiler still needs the separate complete query root.

Proposed integration

PlanningWorkload + evidence + capabilities
    -> ASAPPlanner
    -> PlanSpace
    -> lifecycle-aware selection and materialization
    -> SummaryMaintenanceLifecyclePlan
         - complete selected Post-ASAP root
         - maintenance decisions for its summary producers
         - summary-versus-raw recomputation decision
    -> backend physical compiler validates and binds the selected result
         -> PrecomputePlan
         -> QueryPlan
         -> Summary Catalog definitions
    -> installation and execution
Boundary Current implementation Proposed integration
Computation supplied to compiler Complete selected DAG root Complete root retained in the lifecycle-aware result
Lifecycle selection Called from inside physical compilation for extracted producers Completed through Planner's helper before physical compilation
Connecting DAG and maintenance decisions Backend combines the separate call results Compiler consumes their association in the selected result
Compiler responsibility Obtains lifecycle decisions and generates physical plans Validates feasibility and generates physical plans from the selected decisions
Raw recomputation Must be handled by the existing selection/compilation paths Explicitly honor the helper's raw-versus-summary decision

PlanSpace remains Planner's canonical logical output; lifecycle-aware selection is a helper over that output, as described in #445. The compiler still needs physical implementation, schema, storage and installation context. The simplification concerns the computation/lifecycle handoff, not removal of backend responsibilities.

Relative to the current PR #737 design, this is a smaller change: the document already requires the selected DAG plus lifecycle commitments. The proposal makes their handoff one associated result instead of independently supplied artifacts. The PrecomputePlan/QueryPlan split, catalog definitions and state references remain applicable.

The essential condition is that the result retains the complete query root, including readouts and remaining query operations. For multiple queries, preserve query-to-root associations and shared producer identity across the results so compilation does not duplicate maintenance. If raw recomputation is selected, do not unconditionally install summary producers.

This is a target integration proposal, not an interface the current backend already implements. It requires moving the lifecycle-selection call boundary and preserving those associations, not merely changing a parameter type.


## Architecture at a glance

The current `PrecomputePlan.executable_dags` can contain a complete semantic DAG,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pls add a note that this is confusing and must be changed. Related #740

Comment on lines +82 to +91
lifecycle_commitment:
mode: batch_rebuild_from_data_at_rest
rebuild_every: 1m
retain_for: 10m

backend_capabilities_and_evidence:
supported_modes: [batch_rebuild_from_data_at_rest]
supported_algorithms: [kll]
kll_200_state_bytes: 4096
five_minute_rebuild_cpu_ms: 35

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I do not understand these. Is there documentation?

- id: def-api-latency-kll
input: request_latency_seconds
group_by: [service]
range: 5m

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Examples are helpful thank you. What does range mean?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

range: 5m means the logical input window summarized by the KLL. For an evaluation at time T, it includes samples with timestamps in (T - 5m, T], grouped by service.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

So basically it's equal to the size of the time window (tumbling or sliding) used to generate the KLL instances?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

yes

The baseline is merged code, not the completion of open PRs.

| Area | Existing foundation | Consolidation needed |
“Maintenance” is the execution phase that constructs or updates state, including

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

These concepts are also present in ASAPPlanner right? Wondering if these have been described there

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

yes

Comment thread docs/design_docs/asapplanner-integration.md Outdated
A legal target alternative is:
| Output | Responsibility |
| --- | --- |
| Catalog/SDS entries | Summary semantics, materialization identity, schema and state references |

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

How is this catalog actually used?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Catalog stores the semantics / description of SDS (changed less often), the summary store/sketch store stores the SDS instances payloads (changed per instance).

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Catalog is metadata store

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I understand what it is. I am curious, how it is used right now. Is it used right now?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

catalog is being used rn.

How it's being used --

  • Compilation: The physical compiler constructs SummaryCatalog from the selected precompute configurations, binds PrecomputePlan to it, and validates QueryPlan against it. Compiler code
  • Runtime installation: When the startup physical plan contains a catalog, the data plane installs it into SketchStore. Startup code
  • Query execution: One concrete consumer is MetricsQL per-series readout: it looks up the bound definition in the catalog, resolves its data descriptor, and restores the metric’s name
    label. If that metadata is unavailable, that path requests fallback. Readout code

The current catalog contains summary_descriptors, data_descriptors, and a materializations map. The simplified catalog described in this PR is a proposed change to that existing
representation—not the introduction of a previously unused catalog. Current type

Comment thread docs/design_docs/summary-catalog-sds-architecture.md Outdated
| Object | Meaning | Changes when |
| --- | --- | --- |
| `SummaryDefinition` | Canonical input, operation, grouping, time semantics, algorithm and parameters | Summary semantics change |
| `Materialization` | An installed decision to produce a definition with one state contract | Plan generation or physical contract changes |

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Confused by this. I dont understand the meaning. Also, how is this "Materialization" related to the discussion at #736 (comment)

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

  • SummaryDefinition: summarize request_latency_seconds by service over five minutes using KLL with k=200.
  • SummaryStateInstance: one concrete stored result from that producer, such as the summary for service=api covering (12:00, 12:05].

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Agreed—the standalone catalog Materialization object was an over-abstraction. I have removed it from the proposed design and examples rather than introducing another name for the same layer.

Its information now belongs to the objects that use it:

Information Owner
Summary meaning, state family, algorithm parameters SummaryDefinition in the catalog
Plan version The installed plan bundle; persisted instances also record it for recovery validation
Connection between writer and readers A compiler-assigned state_slot_id in their StateReference, scoped to the plan version
Schema/encoding and physical partition rules PrecomputePlan writer configuration and matching QueryPlan reader configuration
Authorized writer and maintenance policy The PrecomputePlan binding and selected producer lifecycle
Physical-to-semantic provenance The compiler's provenance mapping
Actual partition, coverage, readiness, location and payload format SummaryStateInstance runtime metadata; payload bytes remain in the summary store

The resulting flow is simply:

PrecomputePlan: Build KLL -> Write state slot S
QueryPlan:     Read state slot S -> Estimate p99
Catalog:       SummaryDefinition referenced by both bindings

The state slot is only a join key within a plan version, not a new catalog object with an independent lifecycle. The compiler emits both bindings from one decision and validates their agreement before installation.

Regarding #736: BackendNodeBinding::Materialization remains the existing node-placement marker meaning “store this node's output.” It does not require a separate catalog Materialization object. I have made that distinction explicit.

The docs now remove the proposed materializations collection, update the diagrams and examples, and describe migration through versioned adapters: map existing stored-output identities to slots, retain payload locators, and preserve the format/partition constraints in reader/writer bindings. Existing persisted IDs and wire fields must not be silently renamed or reinterpreted. This PR remains a design-document change; runtime migration is follow-up implementation work.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Okay thanks

| `SummaryDefinition` | Canonical input, operation, grouping, time semantics, algorithm and parameters | Summary semantics change |
| `Materialization` | An installed decision to produce a definition with one state contract | Plan generation or physical contract changes |
| `SummaryStateInstance` | One stored partition, such as a series/pane or completed aggregate | Runtime creates or replaces payload state |
| `StateReference` | A typed plan reference to permitted materialized state | A compiled reader/writer binding changes |

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I did not understand the purpose of this.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

statereference is how plan can reference the summarystateinstance.

| Selected post-ASAP DAG | Planner-selected computation graph, including summary producers, shared dependencies and query readouts. Called the “semantic DAG” in earlier discussion. |
| Summary producer | An operation or subgraph that builds summary state. Multiple queries may share its stored output. |
| `SummaryMaintenanceLifecyclePlan` | Planner result associating a post-ASAP root with deployment decisions for its unique reachable summary producers, plus workload and costing context. |
| Lifecycle commitment | Selected maintenance promise for one producer, with its scheduling and retention binding. A deployment's `SummaryMaintenanceLifecycleGuarantee` carries the Planner-level commitment. |

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Still confused on this. Why do we need this concept?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Good question. I used “lifecycle commitment” only as shorthand for the selected, executable maintenance decision for one summary producer: the selected SummaryMaintenanceLifecycleGuarantee together with its concrete schedule and retention binding. This distinguishes the per-producer decision from the SummaryMaintenanceLifecyclePlan, which contains the DAG, all producer deployments, alternatives, and workload/cost context. It is not a new model or API type. If the shorthand obscures that distinction, I can use “selected deployment guarantee and schedule/retention” throughout instead.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Updated the glossary, integration design, migration plan, and conceptual YAML example to use “selected deployment guarantee and schedule/retention” instead. This is now commit ac855b2.

window framework. The plan also carries workload demand and costing context.
Thus the lifecycle plan already refers to the computation DAG; it is not a
separate query representation, nor is one whole lifecycle plan required per
producer. A missing guarantee is not an executable maintenance commitment.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Did not understand last sentence on missing guarantee

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Agreed; that sentence was too compressed. I rewrote it in ea3226c. In the Planner type, a deployment’s guarantee is optional (None when no alternative is selectable). For that producer, the backend has no selected maintenance mode/schedule to implement, so it must not invent one. If a query needs that stored state, it must use an explicit supported fallback or the plan must be rejected.

physical plans plus catalog bindings. A shared producer is maintained once for
all compatible consumers.

### Lifecycle commitment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Very confused by this. Can you give an example situation where if I do not have this concept, there is some issue with correctness or performance or something else?

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I saw the compiler input example, but still do not understand this concept.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

You are right that the old “lifecycle commitment” wording made this sound like a separate concept. I removed that term and added a concrete example in 118e363. Suppose two queries share a five-minute KLL summary and ask for a result every minute. The selected deployment says to rebuild at each minute boundary and retain completed snapshots for ten minutes. The DAG identifies the shared computation, but by itself does not specify those maintenance decisions. If the backend independently rebuilds only every five minutes, four of five requested endpoints have no matching state; reusing an older snapshot as if it covered the requested interval would change the answer. So the need is to carry and validate the existing Planner-selected guarantee and schedule/retention, not introduce a new model or type.

Comment thread docs/design_docs/asapplanner-migration-plan.md
There is no separate catalog `Materialization` object.

The control plane reconciles two explicitly separate views:
```mermaid

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

For the edges to SummaryDefinition catalog, which are reads and which are writes?
Related, are you saying that if I think of an SDS, the SummaryStore only stores SDS payload, while other metadata stays in the SummaryDefinition metadata catalog?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Good catch; the arrows were ambiguous. I labeled them in 1effeb5. The compiler/catalog authority writes the SummaryDefinition when installing a plan. The PrecomputePlan and QueryPlan bindings reference that definition; those edges mean catalog reads for validation at installation, not runtime catalog writes or serving-time searches. At runtime, PrecomputePlan writes encoded summary payload bytes to the Summary Store and publishes partition/coverage/format/readiness/location as SummaryStateInstance metadata in the runtime inventory. QueryPlan resolves a ready instance from that inventory and reads its payload. So SDS is the contract across the definition catalog, installed plan bindings, runtime instance metadata, and payload store; the other metadata is not all in SummaryDefinition.

the plan. The edges from both plans to that catalog are definition references
validated by catalog reads at installation, not runtime writes or serving-time
catalog searches. At runtime, PrecomputePlan writes summary payload bytes to
the store and publishes each instance's metadata to the inventory. QueryPlan

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Better names for these might be -- SummaryPayloadStore and SummaryMetadataStore ?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

good idea

the store and publishes each instance's metadata to the inventory. QueryPlan
checks the inventory for a ready matching instance, then reads its payload from
the store. SDS describes this combined contract; its metadata is not all stored
in the Summary Catalog. Definition semantics live in the catalog,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This is confusing. Later, you describe the Catalog as storing SummaryDefinition, not metadata?


query_plans:
q50:
read_state: &shared_read

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Meaning of &shared_read and *shared_read ?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Correction: those were YAML anchor/alias syntax and distracted from the design. I removed both and expanded the complete read_state binding under q50 and q99.

| Object | Meaning | Changes when |
| --- | --- | --- |
| `SummaryDefinition` | Canonical input, operation, grouping, time semantics, algorithm and parameters | Summary semantics change |
| `SummaryStateInstance` | One stored partition, such as a series/pane or completed aggregate | Runtime publishes a new or replacement instance |

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Does this include both payload and metadata?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Correction: yes, conceptually. SummaryStateInstance means one complete logical SummaryStore entry: instance metadata plus its associated payload. An implementation may keep the bytes separately internally, but that does not create another architecture component.

encoding: kll-binary-v1
partition: {service: api, window_end: '12:05'}
coverage: {start_exclusive: '12:00', end_inclusive: '12:05'}
location: opaque-store-locator

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

what is this?

encoding: kll-binary-v1
partition_by: [service, window_end]

state_instances:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

What does it mean for a plan to have state_instances? This YAML block seems to describe the metadata of a particular SDS instance.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Correction: runtime instances are not part of the plan. The example now has separate installed_plan and runtime_summary_store sections. The latter is explicitly observed runtime data produced after the PrecomputePlan writes state.

state_instances:
- id: state-api-1205
plan_version: 42
state_slot_id: latency-kll

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

what is slot?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Correction: I renamed state_slot_id to stored_output_id and defined it as the plan-version-scoped binding ID for one persisted PrecomputePlan DAG output. The writer and QueryPlan readers use it to name the same output; it is not a memory slot or storage object.

Comment on lines +225 to +229
| `Desired` | Installed plan | The plan requires state for this slot and coverage |
| `Building` | Runtime inventory | Required state is being produced or recovered |
| `Ready` | Runtime inventory | Required schema and coverage are available |
| `Draining` | Runtime inventory | New work has stopped while existing use completes |
| `Retired` | Runtime inventory | New reads are prohibited; safe reclamation may follow |

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

What's the usecase for these?

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The five-phase table mixed plan intent, build progress, and storage retirement without a current SDS use case. I replaced it with the one decision the DAG path needs now: a QueryPlan read is eligible only when its bound instance exists, the payload is committed, and version, definition, format, partition, and coverage match. Building/draining/retirement remain existing runtime concerns rather than new SDS states.

@zzylol
zzylol changed the base branch from refactor/planner-selection-api to main September 28, 2026 13:47
@zzylol
zzylol force-pushed the docs/physical-plan-design branch from b0d77ce to 5f1eebf Compare September 28, 2026 13:47
@zzylol

zzylol commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor Author

A key idea of this SDS design is that its semantic definition describes the whole computation sub-DAG being precomputed, up to the persisted output—not just the final summary operator. For example, KLL(latency) and KLL(log(latency)) must have different semantic definitions even if their stored state formats are identical.

This is related to how we define edge schemas in ASAPPlanner's Pre-ASAP and Post-ASAP DAGs. There is an open design choice:

  • Accumulated semantics: the description on an edge carries the semantics of the upstream sub-DAG producing that value.
  • Local schema: the edge describes only the value passed between its two adjacent nodes; the upstream computation semantics are represented by the DAG and collected when constructing an SDS definition.

The SDS requirement is to preserve the semantic meaning of the precomputed sub-DAG. Whether that meaning should be accumulated into edge descriptions or derived from the DAG separately from local edge schemas remains an open design discussion.

@milindsrivastava1997 added the discussion comment here.

@zzylol
zzylol merged commit b7c0f08 into main Sep 28, 2026
1 check passed
zzylol added a commit that referenced this pull request Sep 28, 2026
…tion

Level 1 declared one expected plan per query while its stated purpose was to
enumerate the candidate inventory before pricing. A single expected shape is a
selection assertion in structural clothing, and four of the ten queries carried
no expectation at all.

Replace it with a membership contract. Each query declares the shapes the
inventory must expose and how this fixture resolves them: MustBind, or
MustReject with a named policy reason. The heap families Planner exposes for
spatial-topk and topk-rate are now required to be present *and* refused, which
is a positive assertion about the inventory rather than silence about it.

Separate policy refusals from defects. A rejection that is neither a declared
policy reason nor a recorded defect now fails the test. Three candidates fail
with "Planner logical fragment does not match any original query subtree" and
one with a schema incompatibility; main is green on these paths, so each is a
regression introduced inside the #737 -> #728 stack. They are listed in
KNOWN_BINDING_DEFECTS with exact occurrence counts, so a new instance fails and
a fixed one forces its entry to be deleted. All must be gone before #775
installs and executes these candidates.

Stop committing derived plans. candidates/ and ensembles/ were ~94k lines, are
fully reproducible from the test, and every identity in them changes when the
Planner pin moves; CI already uploads them on every run. Only the per-query
admission reports stay in the tree, and they now reproduce byte-for-byte from
this build. Historical synthetic-price exports move to archive/, and the
README's hand-copied Planner revision is dropped in favour of reading the pin
from Cargo.toml, since the copied one had gone stale.

The sort-key contract test now compiles the query instead of reading a
committed export, so it no longer depends on plans from an older Planner.

A certified companion fixture for the heap candidates is NOT included: the
issue-754 generator cannot supply one. Scoped evidence needs
topk_selected_lower_bound > topk_excluded_upper_bound, and under the
generator's own per-series domain the third- and fourth-ranked series overlap
in both groups, while these queries evaluate in real_time scope so the bound
must hold across the whole validity window. That needs a generator with
non-overlapping per-series domains.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
zzylol added a commit that referenced this pull request Sep 28, 2026
An earlier revision claimed every recorded defect was a regression introduced
inside the #737 -> #728 stack, on the grounds that main is green on these
paths. Diffing residual_nodes against main does not support that: the function
is substantively identical there, and the stack changed only how the accuracy
target is derived and the wording of the error message.

main is green because its tests never plan these queries, not because the code
is correct. That is the more worrying reading, so it should be the recorded one.

The fragment mismatch is fixed on main by #781: a context-derived leaf schema
was compared for equality against an isolated re-parse that cannot reproduce
it, and open schemas are now compared by containment.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
semantic_format_version: <version>
semantics: <canonical-typed-description>
output: <described-output>
```

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This might be the data structure but it's not exposing the semantics of a definition i.e. the stuff you wrote above like source and filter semantics, value expressions, etc.

Comment on lines +138 to +139
tenant-A/requests ── bound to endpoint east ── definition D_A
tenant-A/requests ── bound to endpoint west ── definition D_A

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

what is endpoint east and west?

Comment on lines +150 to +154
The authority resolving a source preserves dataset identity across relocation and
assigns a different identity when the logical dataset changes. Installation checks
that the concrete source realizes the identity in the definition. Equal definitions
do not grant cross-tenant or cross-deployment authorization. Future discovery must
match dataset identity as well as expression semantics.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

didnt understand this

group_key: {service: api}
window: {start_exclusive: '12:00', end_inclusive: '12:01'}
definition_id: <KLL-latency-definition>
revision: <input-revision>

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

what is revision

Comment on lines +226 to +228
stored_output_id
= Which authorized deployed output does this state belong to?
```

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

didnt get this

```

Both outputs have identical semantics, but hot may be the active serving output
while rebuild is still being validated. Even adding plan version to definition

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

rebuild is still being validated

didn't get this

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants